Papers with semi-supervised approach
Bilingual Lexicon Induction with Semi-supervision in Non-Isometric Embedding Spaces (P19-1)
Copied to clipboard
| Challenge: | Recent work on bilingual lexicon induction (BLI) relies on an assumption about the isometry of two embedding spaces. |
| Approach: | They propose a semi-supervised approach that relaxes the isometric assumption while leveraging limited aligned bilingual lexicons and a larger set of unaligned word embeddings. |
| Outcome: | The proposed method obtains state-of-the-art results on 15 of 18 language pairs on the MUSE dataset and does particularly well when the embedding spaces don’t appear isometric. |
Consistent Text Categorization using Data Augmentation in e-Commerce (2023.acl-industry)
Copied to clipboard
| Challenge: | Upon closer inspection, we found inconsistencies in the labeling of similar items. |
| Approach: | They propose to improve an existing product categorization model that takes a product title as input and outputs the most suitable category out of thousands of available candidates. |
| Outcome: | The proposed model is based on a product title and outputs the most suitable category out of thousands of available candidates. |
Disentangled Learning of Stance and Aspect Topics for Vaccine Attitude Detection in Social Media (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing approaches to detect vaccine attitudes on social media require abundant annotations and pre-defined aspect categories. |
| Approach: | They propose a semi-supervised approach to detect vaccine attitudes on social media . they use an autoencoding architecture to learn from unlabelled data the topical information of the domain . |
| Outcome: | The proposed model outperforms existing aspect-based models on stance detection and tweet clustering. |
With More Contexts Comes Better Performance: Contextualized Sense Embeddings for All-Round Word Sense Disambiguation (2020.emnlp-main)
Copied to clipboard
| Challenge: | Contextualized word embeddings have been used effectively across several tasks in Natural Language Processing, but it is difficult to link them to structured sources of knowledge. |
| Approach: | They propose a semi-supervised approach to producing sense embeddings for the lexical meanings within a lexicon that is comparable to that of contextualized word vectors. |
| Outcome: | The proposed approach outperforms state-of-the-art models in the English Word Sense Disambiguation task and in the multilingual one while training on sense-annotated data in English only. |
Semi-Supervised Disfluency Detection (C18-1)
Copied to clipboard
| Challenge: | Detecting disfluency can be difficult because of the flexible nature of reparandum structure and the lack of a nested structure. |
| Approach: | They propose a semi-supervised approach which extracts hidden features from self-attention without any Recurrent Neural Network (RNN) or Convolutional Neural Net (CNN). |
| Outcome: | The proposed approach improves over baselines by using unlabelled data . identifying and removing non-fluent factors would help to improve spontaneous speech quality . |
Plausible Extractive Rationalization through Semi-Supervised Entailment Signal (2024.findings-acl)
Copied to clipboard
| Challenge: | Abstract: Large language models are gaining widespread adoption in natural language processing tasks. |
| Approach: | They propose a semi-supervised approach to optimize for plausibility of extracted rationales by using a pre-trained natural language inference model and a supervised NLI predictor. |
| Outcome: | The proposed model outperforms unsupervised models by > 100% on a ERASER dataset. |
Not Far Away, Not So Close: Sample Efficient Nearest Neighbour Data Augmentation via MiniMax (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing kNN-based augmentation techniques blindly incorporate all samples, but MiniMax-kNN uses a subset of augmented samples to maximize KL-divergence between teacher and student models. |
| Approach: | They propose a semi-supervised approach to augmented data augmentation using kNN. |
| Outcome: | The proposed method outperforms existing kNN-based augmentation techniques on several classification tasks and requires fewer augmented examples and less computation to achieve superior performance. |
Low-Confidence Gold: Refining Low-Confidence Samples for Efficient Instruction Tuning (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Low-Confidence Gold (LCG) is a new filtering framework for Large Language Models that curates high-quality subsets while preserving data diversity. |
| Approach: | They propose a new filtering framework that employs centroid-based clustering and confidence-guided selection for identifying valuable instruction pairs. |
| Outcome: | The proposed framework improves performance on a subset of 6K samples while maintaining data diversity. |
Self-Training for Sample-Efficient Active Learning for Text Classification with Pre-Trained Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to train models without labeled data are lacking in supervised tasks . a lack of labeles is the main obstacle to real-world applications . |
| Approach: | They propose a semi-supervised approach that uses a model to obtain pseudo-labels for unlabeled data. |
| Outcome: | The proposed method outperforms the reproduced methods on four text classification benchmarks. |
Semi-supervised multimodal coreference resolution in image narrations (2023.emnlp-main)
Copied to clipboard
| Challenge: | a semi-supervised approach is used to resolve multimodal coreferences and narrative grounding in a multimodal context. |
| Approach: | They propose a semi-supervised approach that utilizes image-narration pairs to resolve coreferences and narrative grounding in a multimodal context. |
| Outcome: | The proposed approach outperforms baselines quantitatively and qualitatively for coreference resolution and narrative grounding in a multimodal context. |
ThaiLMCut: Unsupervised Pretraining for Thai Word Segmentation (2020.lrec-1)
Copied to clipboard
Suteera Seeha, Ivan Bilan, Liliana Mamani Sanchez, Johannes Huber, Michael Matuschek, Hinrich Schütze
| Challenge: | ThaiLMCut is a semi-supervised word segmentation model for word segmenting in Thai . it uses a bi-directional character language model to leverage useful linguistic knowledge from unlabeled data. |
| Approach: | They propose a semi-supervised approach to Thai word segmentation using a character language model. |
| Outcome: | The proposed approach outperforms state-of-the-art models on the benchmark InterBEST2009. |